Tag
3 articles
This article explains how to fine-tune the Qwen3 language model using Low-Rank Adaptation (LoRA) and NVIDIA NeMo AutoModel in a single-GPU Google Colab environment, focusing on parameter-efficient training techniques and automated workflows.
Xiaomi's MiMo team, with TileRT, has achieved over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node, marking a significant leap in LLM inference performance.
AutoKernel is an open-source framework that uses autonomous LLM agents to automate GPU kernel optimization for PyTorch models, significantly reducing manual effort and improving performance.